Papers with visual questions
Finding the Evidence: Localization-aware Answer Prediction for Text Visual Question Answering (2020.coling-main)
Copied to clipboard
| Challenge: | Existing text VQA systems generate an answer by selecting from optical character recognition (OCR) texts or a fixed vocabulary. |
| Approach: | They propose a localization-aware answer prediction network that generates the answer and predicts a bounding box as evidence of the generated answer. |
| Outcome: | The proposed network outperforms existing methods on three benchmark datasets for the text VQA task by a noticeable margin. |
Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQA (2023.acl-long)
Copied to clipboard
| Challenge: | Visual question answering models seek to answer questions about images . ambiguity can exist at all levels of linguistic analysis, but disagreements can be difficult to detect and resolve . |
| Approach: | They develop a question-generation model which integrates group information without supervision and uses a dataset of ambiguous examples to annotate answers. |
| Outcome: | The proposed model can integrate answer group information without supervision and is able to fill knowledge gaps and convey requests. |
Answering Cross-Dimensional Geometric Visual Questions by Multi-constraint Spatial Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for solving complex visual questions are limited in their ability to represent in a cross-dimensional space. |
| Approach: | They propose a method that can answer complex visual questions using cross-dimensional reasoning. |
| Outcome: | The proposed method can answer complex visual questions in 2D to 3D space with great application value. |